Parallel Experiments | Telegram Webview: LinghaoCh/938 -

Telegram Group & Telegram Channel

Parallel Experiments

https://arxiv.org/abs/2305.18290 #llm #ai

今天深入学习了 DPO，再次感叹扎实的数学功底对 AI/ML Research 的重要性……

原始的 RLHF 是用 pairwise human preference data（A 和 B 哪个更好）去训练一个 reward model，然后用 RL 来训练主 policy model，objective 是 minimize negative log likelihood + regularization（比如 PPO 就是通过新旧 policy 之间的 KL Divergence 来做 regularization）。这样的缺点在于 RL 是出了名的难搞，而且还需要一个 critic model 来预测 reward，使得整个系统的复杂性很高。

DPO 的思路是，观察到 RLHF 的 objective 本质上是 minimize loss over (latent) reward function，通过一番 reparameterization 等数学推导，重新设计了一个 minimize loss over policy 的 objective，绕过了中间这个 reward model，让 gradient update 直接增加 policy model 生成 winner response 的概率并降低 loser response 的概率，大幅简化了流程。

拓展阅读：
- KTO: 更进一步，不需要 pairwise comparison，只用对 individual example 的 upvote/downvote 也可以学习到 preference。
- IPO: 解决 DPO 容易 overfit 的问题。

Direct Preference Optimization: Your Language Model is Secretly a...

While large-scale unsupervised language models (LMs) learn broad world knowledge and some reasoning skills, achieving precise control of their behavior is difficult due to the completely...

www.tg-me.com/nl/Parallel Experiments/com.LinghaoCh/938

1.3K viewsLinghao Zhang, edited Apr 19 at 05:31

tg-me.com/LinghaoCh/938

Create: 2025-04-19
Last Update: 2025-05-31 14:19:03

https://arxiv.org/abs/2305.18290 #llm #ai

今天深入学习了 DPO，再次感叹扎实的数学功底对 AI/ML Research 的重要性……

原始的 RLHF 是用 pairwise human preference data（A 和 B 哪个更好）去训练一个 reward model，然后用 RL 来训练主 policy model，objective 是 minimize negative log likelihood + regularization（比如 PPO 就是通过新旧 policy 之间的 KL Divergence 来做 regularization）。这样的缺点在于 RL 是出了名的难搞，而且还需要一个 critic model 来预测 reward，使得整个系统的复杂性很高。

DPO 的思路是，观察到 RLHF 的 objective 本质上是 minimize loss over (latent) reward function，通过一番 reparameterization 等数学推导，重新设计了一个 minimize loss over policy 的 objective，绕过了中间这个 reward model，让 gradient update 直接增加 policy model 生成 winner response 的概率并降低 loser response 的概率，大幅简化了流程。

拓展阅读：
- KTO: 更进一步，不需要 pairwise comparison，只用对 individual example 的 upvote/downvote 也可以学习到 preference。
- IPO: 解决 DPO 容易 overfit 的问题。

BY Parallel Experiments

Share with your friend now:
tg-me.com/LinghaoCh/938

Open in Telegram

Parallel Experiments Telegram | DID YOU KNOW?

Date: 2025-05-31| Parallel Experiments

If riding a bucking bronco is your idea of fun, you’re going to love what the stock market has in store. Consider this past week’s ride a preview.The week’s action didn’t look like much, if you didn’t know better. The Dow Jones Industrial Average rose 213.12 points or 0.6%, while the S&P 500 advanced 0.5%, and the Nasdaq Composite ended little changed.

How to Invest in Bitcoin?

Like a stock, you can buy and hold Bitcoin as an investment. You can even now do so in special retirement accounts called Bitcoin IRAs. No matter where you choose to hold your Bitcoin, people’s philosophies on how to invest it vary: Some buy and hold long term, some buy and aim to sell after a price rally, and others bet on its price decreasing. Bitcoin’s price over time has experienced big price swings, going as low as $5,165 and as high as $28,990 in 2020 alone. “I think in some places, people might be using Bitcoin to pay for things, but the truth is that it’s an asset that looks like it’s going to be increasing in value relatively quickly for some time,” Marquez says. “So why would you sell something that’s going to be worth so much more next year than it is today? The majority of people that hold it are long-term investors.”

Parallel Experiments from nl

Telegram Parallel Experiments
FROM USA